Skip to content

feat(desktop): live conversation token rate (tok/s) - #416

Open
GISWLH wants to merge 4 commits into
vastsa:mainfrom
GISWLH:feat/conversation-token-rate
Open

GISWLH wants to merge 4 commits into
vastsa:mainfrom
GISWLH:feat/conversation-token-rate

Conversation

@GISWLH

@GISWLH GISWLH commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Fixes #394

Summary

  • Show a live sliding-window output rate (tok/s) on the transcript stream-health strip while a turn is running (working / run-activity indicators, plus a compact chip while answer tokens stream).
  • Prefer provider outputTokens mid-stream; otherwise estimate from visible thinking+answer text (four code points ≈ one token) and label as approximate (≈).
  • Context inspector Generation speed remains the completed-turn snapshot (D428).

Algorithm

  • Cumulative samples on a stable ~250ms interval (useLiveTokenRate keeps content/thinking/outputTokens in refs so stream deltas do not tear down the timer).
  • Rate = rounded Δtokens / Δt over a 2s sliding window (min duration 200ms).
  • TTFT: until some positive output has been seen, rate stays undefined (chip hidden)—never a false 0 tok/s.
  • Stall: after positive output, a drained window reports 0 so reconnect/stall is visible.
  • Estimate → provider handoff: if the source flips to provider usage and the provider count is lower than the estimate, reset the sample window so monotonic clamping cannot freeze a hard 0 without ≈.

Review fixes (this push)

  • a11y: ticking rate is aria-hidden / decorative (mirrors working-elapsed); streaming-only chip is not a polite live region.
  • Guard mounts with tokensPerSecond !== undefined (not always-truthy {tokenRate ? …}).
  • zh-CN mirrors for component-spec, interaction-patterns §3.2a, E2E plan, decisions-log / D428.
  • FR live copy uses jetons/s.

Tests

  • pnpm typecheck / pnpm lint in apps/desktop — green
  • node --test test/streaming-token-rate.test.mjs test/active-turn-surface.test.mjs — green (TTFT undefined, estimate→provider reset, hook interval deps, a11y/guard contracts)

Surface a sliding-window tok/s reading on the transcript stream-health
strip so users can tell whether the model stream is healthy or stalled.
Prefer provider output usage when available; otherwise estimate from
visible text and label as approximate. Fixes vastsa#394.
Keep the chip hidden until positive output (no false TTFT 0), reset the
sample window on estimate→provider handoff when provider is lower, sample
on a stable 250ms interval via input refs, and treat the ticking rate as
decorative for a11y. Mirror D428 in zh-CN specs and use FR jetons/s.
Resolve conflicts with main:
- ChatTranscript/ActivityGroup: re-apply the live tok/s chip on top of the
  always-mounted runtime status lane and the extracted ActivityItems; drop
  the separate streaming-only indicator since the lane now stays mounted
  for the whole running turn.
- Spec/E2E/decision-log: keep upstream entries; renumber this change's
  decision from D428 to D637 (D428 is now taken upstream).
- i18n: add the live throughput keys to the new pt-BR locale.
Resolve decision-log conflicts: keep upstream entries; renumber this change's
decision from D637 to D639 (D637 and D638 are now taken upstream).
@GISWLH

GISWLH commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Merged the latest main into this branch to resolve the conflicts (merge commit rather than rebase, since CONTRIBUTING/AGENTS.md say not to force-push contributor branches, so this is a plain fast-forward push).

What conflicted and how it was resolved

  • ChatTranscript.tsx / ActivityGroup.tsx: main moved the activity rendering into ActivityItems / ProcessActivityGroup and wrapped the tail indicators in an always-mounted transcript-runtime-status lane ([Bug] “等待模型响应”等临时状态出现/消失时,会话内容上下跳动 #323). I took main's versions and re-applied the tok/s changes on top: LiveTokenRateLabel, the optional tokenRate prop on WorkingIndicator and RunActivityIndicator, and the useLiveTokenRate wiring. The lane now stays mounted for the whole running turn, so the separate StreamingTokenRateIndicator is no longer needed. I dropped it, along with its CSS rule and test assertions, and updated the spec/E2E wording to match.
  • Decision log and E2E plan (en + zh-CN): kept all the new upstream entries and re-added E2E-CHAT-live-token-rate-shows-during-stream to the Conversation & stream row. D428 is now taken upstream ("A tooltip never outlives its trigger"), so this change's decision is renumbered to D639 (D637 and D638 have since been taken upstream too), and its references in the specs are updated.
  • Not flagged by git, but needed: main added a pt-BR locale, so packages/i18n failed to typecheck without usageLiveThroughput / usageLiveThroughputEstimated. I added both keys to pt-BR.

Checks run locally (Node 22.23, pnpm 12.8.1)

  • pnpm typecheck: pass
  • pnpm lint: pass (biome and style tokens)
  • pnpm -r --if-present test: shared, plugin-sdk, i18n, host-runtime, voice-runtime, plugin-devkit and pi-host all pass. apps/desktop passes 3344 of 3356; the other failures are the same ones that fail on an untouched upstream/main checkout (chat-error-message.test.mjs "network failures show the transport errno…" and the plugin-websocket.test.mjs tests). They look environment-related and are unrelated to this PR. active-turn-surface and streaming-token-rate pass.
  • cargo checks not run, since this change is renderer-only.
  • E2E suites not run (no Electron/display here), so the E2E gate for main still needs to be run by a maintainer or in a capable environment.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]对话token速率显示

1 participant